Papers with data-centric analysis

2 papers
Early Guessing for Dialect Identification (2022.findings-emnlp)

Copied to clipboard

Challenge: Current research on dialect identification is model-centric, focusing on performance.
Approach: They propose a data-centric approach to find the shortest input needed to make a plausible guess.
Outcome: The proposed method generalizes across dialects and datasets with two shortening criteria.
Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active Exploration (2023.emnlp-main)

Copied to clipboard

Challenge: Existing data for instruction-tuning are inadequate for a wide range of tasks, limiting the scope for nuanced comprehension and interactions within these domains.
Approach: They propose to use Large Language Models to explore a multitude of variations or possibilities to improve instruction-tuning data by active exploration.
Outcome: The proposed approach improves domain-specific instruction coverage and shows significant improvements over baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations